Papers with Reddit communities
Dreaddit: A Reddit Dataset for Stress Analysis in Social Media (D19-62)
Copied to clipboard
| Challenge: | Existing computational studies on stress only focus on domains such as speech or Twitter . a corpus of social media text is used to identify stress . |
| Approach: | They propose a text corpus of lengthy social media data for detecting stress . they use 190K posts from five different categories of Reddit communities . |
| Outcome: | The proposed corpus of social media data can be used to identify stress . it includes 190K posts from five different categories of Reddit communities . |
Characterizing English Variation across Social Media Communities with BERT (2021.tacl-1)
Copied to clipboard
| Challenge: | Existing studies characterizing language variation across Internet social groups have focused on the types of words used by these groups. |
| Approach: | They extend this study by employing BERT to characterize variation in the senses of words as well, analyzing two months of English comments in 474 Reddit communities. |
| Outcome: | The proposed study analyzes two months of English comments in 474 Reddit communities and ties language variation with community behavior. |
MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language Models (2025.emnlp-main)
Copied to clipboard
Zhengyi Zhao, Shubo Zhang, Yuxi Zhang, Yanxi Zhao, Yifan Zhang, Zezhong Wang, Huimin Wang, Yutian Zhao, Bin Liang, Yefeng Zheng, Binyang Li, Kam-Fai Wong, Xian Wu
| Challenge: | Current approaches focus on isolated meme analysis, either for harmful content detection or standalone interpretation, overlooking a fundamental challenge: the same meme can express different intents depending on its conversational context. |
| Approach: | They propose a benchmark to evaluate how large vision language models understand memes in their original context. |
| Outcome: | The proposed benchmark evaluates how large vision language models understand meme intent in their original context. |
“Are you kidding me?”: Detecting Unpalatable Questions on Reddit (2021.eacl-main)
Copied to clipboard
| Challenge: | Existing methods to detect online abuse focus on the more explicit forms of abuse . existing methods focus on detecting subtler forms of online abuse leaving them unnoticed . |
| Approach: | They propose a task to detect unpalatable questions using reddit data to implement a context-aware dataset and implement 'learning models' they hope future research will address subtle forms of abuse since harm passes unnoticed through existing detection systems. |
| Outcome: | The proposed task is based on a dataset of reddit users and a conversational context. |
Investigating Online Community Engagement through Stancetaking (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Large-scale computational work on stancetaking has explored community similarities in their preferences for stance markers without considering the stance-relevant properties of the contexts in which stance marker use is carried out. |
| Approach: | They propose to use stance-relevant properties of Reddit communities to capture community identity patterns distinct from textual or marker similarity measures. |
| Outcome: | The proposed representations capture community identity patterns distinct from textual or marker similarity measures and relate them to broader inter- and intra-community engagement patterns. |
SLM-Mod: Small Language Models Surpass LLMs at Content Moderation (2025.naacl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) are expensive to query in real-time and do not allow for a community-specific approach to content moderation. |
| Approach: | They propose to use small language models for community-specific content moderation tasks by fine-tuning and evaluating their performance against larger open- and closed-sourced models. |
| Outcome: | The proposed models outperform zero-shot LLMs in content moderation tasks with 11.5% higher accuracy and 25.7% higher recall across all communities. |
STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models can adapt outputs to align with community-specific norms, perspectives and communication styles. |
| Approach: | They propose a benchmark to assess community-specific steering using contrasting reddit communities. |
| Outcome: | STEER-BENCH assesses how well large language models understand community-specific instructions, their resilience to adversarial steering attempts, and their ability to accurately represent cultural and ideological perspectives. |
ValueScope: Unveiling Implicit Norms and Values via Return Potential Model of Social Interactions (2024.findings-emnlp)
Copied to clipboard
Chan Young Park, Shuyue Stella Li, Hayoung Jung, Svitlana Volkova, Tanu Mitra, David Jurgens, Yulia Tsvetkov
| Challenge: | VALUESCOPE is a framework that quantifies social norms and values within online communities. |
| Approach: | They propose a framework that uses language models to quantify social norms and values within online communities. |
| Outcome: | The proposed framework delineates differences in social norms and tracks evolution of norms in online communities and influence of significant external events like the U.S. presidential elections and the emergence of new sub-communities. |
PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media (2026.acl-long)
Copied to clipboard
| Challenge: | Social media are shifting towards community-governed platforms where groups define their own norms. |
| Approach: | They propose a multimodal, multilingual benchmark for detecting 13,371 rule violations across 1,989 Reddit communities . they show that bigger models and increased context provide marginal gains, and universal rules like civility and self-promotion are easier to detect. |
| Outcome: | The proposed model can detect 13,371 rule violations across 1,989 Reddit communities across 2,885 rules in 9 languages. |